539 research outputs found
Large introns in relation to alternative splicing and gene evolution: a case study of Drosophila bruno-3
Background:
Alternative splicing (AS) of maturing mRNA can generate structurally and functionally distinct transcripts from the same gene. Recent bioinformatic analyses of available genome databases inferred a positive correlation between intron length and AS. To study the interplay between intron length and AS empirically and in more detail, we analyzed the diversity of alternatively spliced transcripts (ASTs) in the Drosophila RNA-binding Bruno-3 (Bru-3) gene. This gene was known to encode thirteen exons separated by introns of diverse sizes, ranging from 71 to 41,973 nucleotides in D. melanogaster. Although Bru-3's structure is expected to be conducive to AS, only two ASTs of this gene were previously described.
Results:
Cloning of RT-PCR products of the entire ORF from four species representing three diverged Drosophila lineages provided an evolutionary perspective, high sensitivity, and long-range contiguity of splice choices currently unattainable by high-throughput methods. Consequently, we identified three new exons, a new exon fragment and thirty-three previously unknown ASTs of Bru-3. All exon-skipping events in the gene were mapped to the exons surrounded by introns of at least 800 nucleotides, whereas exons split by introns of less than 250 nucleotides were always spliced contiguously in mRNA. Cases of exon loss and creation during Bru-3 evolution in Drosophila were also localized within large introns. Notably, we identified a true de novo exon gain: exon 8 was created along the lineage of the obscura group from intronic sequence between cryptic splice sites conserved among all Drosophila species surveyed. Exon 8 was included in mature mRNA by the species representing all the major branches of the obscura group. To our knowledge, the origin of exon 8 is the first documented case of exonization of intronic sequence outside vertebrates.
Conclusion:
We found that large introns can promote AS via exon-skipping and exon turnover during evolution likely due to frequent errors in their removal from maturing mRNA. Large introns could be a reservoir of genetic diversity, because they have a greater number of mutable sites than short introns. Taken together, gene structure can constrain and/or promote gene evolution
Recommended from our members
The Single-Nucleotide Resolution Transcriptome of Pseudomonas aeruginosa Grown in Body Temperature
One of the hallmarks of opportunistic pathogens is their ability to adjust and respond to a wide range of environmental and host-associated conditions. The human pathogen Pseudomonas aeruginosa has an ability to thrive in a variety of hosts and cause a range of acute and chronic infections in individuals with impaired host defenses or cystic fibrosis. Here we report an in-depth transcriptional profiling of this organism when grown at host-related temperatures. Using RNA-seq of samples from P. aeruginosa grown at 28°C and 37°C we detected genes preferentially expressed at the body temperature of mammalian hosts, suggesting that they play a role during infection. These temperature-induced genes included the type III secretion system (T3SS) genes and effectors, as well as the genes responsible for phenazines biosynthesis. Using genome-wide transcription start site (TSS) mapping by RNA-seq we were able to accurately define the promoters and cis-acting RNA elements of many genes, and uncovered new genes and previously unrecognized non-coding RNAs directly controlled by the LasR quorum sensing regulator. Overall we identified 165 small RNAs and over 380 cis-antisense RNAs, some of which predicted to perform regulatory functions, and found that non-coding RNAs are preferentially localized in pathogenicity islands and horizontally transferred regions. Our work identifies regulatory features of P. aeruginosa genes whose products play a role in environmental adaption during infection and provides a reference transcriptional landscape for this pathogen
Characteristics of transposable element exonization within human and mouse
Insertion of transposed elements within mammalian genes is thought to be an
important contributor to mammalian evolution and speciation. Insertion of
transposed elements into introns can lead to their activation as alternatively
spliced cassette exons, an event called exonization. Elucidation of the
evolutionary constraints that have shaped fixation of transposed elements
within human and mouse protein coding genes and subsequent exonization is
important for understanding of how the exonization process has affected
transcriptome and proteome complexities. Here we show that exonization of
transposed elements is biased towards the beginning of the coding sequence in
both human and mouse genes. Analysis of single nucleotide polymorphisms (SNPs)
revealed that exonization of transposed elements can be population-specific,
implying that exonizations may enhance divergence and lead to speciation. SNP
density analysis revealed differences between Alu and other transposed
elements. Finally, we identified cases of primate-specific Alu elements that
depend on RNA editing for their exonization. These results shed light on TE
fixation and the exonization process within human and mouse genes.Comment: 11 pages, 4 figure
The Origins, Evolution, and Functional Potential of Alternative Splicing in Vertebrates
Alternative splicing (AS) has the potential to greatly expand the functional repertoire of mammalian transcriptomes. However, few variant transcripts have been characterized functionally, making it difficult to assess the contribution of AS to the generation of phenotypic complexity and to study the evolution of splicing patterns. We have compared the AS of 309 protein-coding genes in the human ENCODE pilot regions against their mouse orthologs in unprecedented detail, utilizing traditional transcriptomic and RNAseq data. The conservation status of every transcript has been investigated, and each functionally categorized as coding (separated into coding sequence [CDS] or nonsense-mediated decay [NMD] linked) or noncoding. In total, 36.7% of human and 19.3% of mouse coding transcripts are species specific, and we observe a 3.6 times excess of human NMD transcripts compared with mouse; in contrast to previous studies, the majority of species-specific AS is unlinked to transposable elements. We observe one conserved CDS variant and one conserved NMD variant per 2.3 and 11.4 genes, respectively. Subsequently, we identify and characterize equivalent AS patterns for 22.9% of these CDS or NMD-linked events in nonmammalian vertebrate genomes, and our data indicate that functional NMD-linked AS is more widespread and ancient than previously thought. Furthermore, although we observe an association between conserved AS and elevated sequence conservation, as previously reported, we emphasize that 30% of conserved AS exons display sequence conservation below the average score for constitutive exons. In conclusion, we demonstrate the value of detailed comparative annotation in generating a comprehensive set of AS transcripts, increasing our understanding of AS evolution in vertebrates. Our data supports a model whereby the acquisition of functional AS has occurred throughout vertebrate evolution and is considered alongside amino acid change as a key mechanism in gene evolution
Comparative transcriptomics of pathogenic and non-pathogenic Listeria species
Comparative RNA-seq analysis of two related pathogenic and non-pathogenic bacterial strains reveals a hidden layer of divergence in the non-coding genome as well as conserved, widespread regulatory structures called ‘Excludons', which mediate regulation through long non-coding antisense RNAs
How the other half lives: CRISPR-Cas's influence on bacteriophages
CRISPR-Cas is a genetic adaptive immune system unique to prokaryotic cells
used to combat phage and plasmid threats. The host cell adapts by incorporating
DNA sequences from invading phages or plasmids into its CRISPR locus as
spacers. These spacers are expressed as mobile surveillance RNAs that direct
CRISPR-associated (Cas) proteins to protect against subsequent attack by the
same phages or plasmids. The threat from mobile genetic elements inevitably
shapes the CRISPR loci of archaea and bacteria, and simultaneously the
CRISPR-Cas immune system drives evolution of these invaders. Here we highlight
our recent work, as well as that of others, that seeks to understand phage
mechanisms of CRISPR-Cas evasion and conditions for population coexistence of
phages with CRISPR-protected prokaryotes.Comment: 24 pages, 8 figure
Deep Transfer Learning on Satellite Imagery Improves Air Quality Estimates in Developing Nations
Urban air pollution is a public health challenge in low- and middle-income countries (LMICs). However, LMICs lack adequate air quality (AQ) monitoring infrastructure. A persistent challenge has been our inability to estimate AQ accurately in LMIC cities, which hinders emergency preparedness and risk mitigation. Deep learning-based models that map satellite imagery to AQ can be built for high-income countries (HICs) with adequate ground data. Here we demonstrate that a scalable approach that adapts deep transfer learning on satellite imagery for AQ can extract meaningful estimates and insights in LMIC cities based on spatiotemporal patterns learned in HIC cities. The approach is demonstrated for Accra in Ghana, Africa, with AQ patterns learned from two US cities, specifically Los Angeles and New York
Spatial particulate fields during highwinds in the imperial valley, California
We examined windblown dust within the Imperial Valley (CA) during strong springtime west-southwesterly (WSW) wind events. Analysis of routine agency meteorological and ambient particulate matter (PM) measurements identified 165 high WSW wind events between March and June 2013 to 2019. The PM concentrations over these days are higher at northern valley monitoring sites, with daily PM mass concentration of particles less than 10 micrometers aerodynamic diameter (PM10) at these sites commonly greater than 100 μg/m3 and reaching around 400 μg/m3, and daily PM mass concentration of particles less than 2.5 micrometers aerodynamic diameter (PM2.5) commonly greater than 20 μg/m3 and reaching around 60 μg/m3. A detailed analysis utilizing 1 km resolution multi-angle implementation of atmospheric correction (MAIAC) aerosol optical depth (AOD), Identifying Violations Affecting Neighborhoods (IVAN) low-cost PM2.5 measurements and 500 m resolution sediment supply fields alongside routine ground PM observations identified an area of high AOD/PM during WSW events spanning the northwestern valley encompassing the Brawley/Westmorland through the Niland area. This area shows up most clearly once the average PM10 at northern valley routine sites during WSW events exceeds 100 μg/m3. The area is consistent with high soil sediment supply in the northwestern valley and upwind desert, suggesting local sources are primarily responsible. On the basis of this study, MAIAC AOD appears able to identify localized high PM areas during windblown dust events provided the PM levels are high enough. The use of the IVAN data in this study illustrates how a citizen science effort to collect more spatially refined air quality concentration data can help pinpoint episodic pollution patterns and possible sources important for PM exposure and adverse health effects
Alu-Alu Recombination Underlying the First Large Genomic Deletion in GlcNAc-Phosphotransferase Alpha/Beta (GNPTAB) Gene in a MLII Alpha/Beta Patient
Mucolipidosis type II α/β is a severe, autosomal recessive lysosomal storage disorder, caused by a defect in the GNPTAB gene that codes for the α/β subunits of the GlcNAc-phosphotransferase. To date, over 100 different mutations have been identified in MLII α/β patients, but no large deletions have been reported. Here we present the first case of a large homozygous intragenic GNPTAB gene deletion (c.3435-386_3602 + 343del897) encompassing exon 19, identified in a ML II α/β patient. Long-range PCR and sequencing methodologies were used to refine the characterization of this rearrangement, leading to the identification of a 21 bp repetitive motif in introns 18 and 19. Further analysis revealed that both the 5' and 3' breakpoints were located within highly homologous Alu elements (Alu-Sz in intron 18 and Alu-Sq2, in intron 19), suggesting that this deletion has probably resulted from Alu-Alu unequal homologous recombination. RT-PCR methods were used to further evaluate the consequences of the alteration for the processing of the mutant pre mRNA GNPTAB, revealing the production of three abnormal transcripts: one without exon 19 (p.Lys1146_Trp1201del); another with an additional loss of exon 20 (p.Arg1145Serfs*2), and a third in which exon 19 was substituted by a pseudoexon inclusion consisting of a 62 bp fragment from intron 18 (p.Arg1145Serfs*16). Interestingly, this 62 bp fragment corresponds to the Alu-Sz element integrated in intron 18.This represents the first description of a large deletion identified in the GNPTAB gene and contributes to enrich the knowledge on the molecular mechanisms underlying causative mutations in ML II.This work was supported by FCT - project PIC/IC/83252/2007 (http://alfa.fct.mctes.pt/). Coutinho MF and Quental S received grants from the FCT (SFRH/BD/48103/2008; SFRH/BPD/64025/2009)
- …